Papers with medical question answering datasets
Enhancing Healthcare LLM Trust with Atypical Presentations Recalibration (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for eliciting and calibrating large language models have focused on general reasoning datasets, yielding only modest improvements. |
| Approach: | They propose a method which leverages atypical presentations to adjust model confidence estimates. |
| Outcome: | The proposed method reduces calibration errors by approximately 60% on three medical question answering datasets and outperforms existing methods such as vanilla verbalized confidence, CoT verbalised confidence and others. |